Papers with critic-free RL framework
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | extending grouping-based methods to agentic reasoning presents unique challenges . frequent environment interactions and tool invocations render intra-group advantage estimation unstable . |
| Approach: | They propose a grouping-based method that uses a single round of rollouts to stabilize advantage estimation. |
| Outcome: | a new RL framework outperforms grouping-based methods in retrieval tasks and advanced mathematical reasoning benchmarks. |